Papers by David R. Mortensen
A Hmong Corpus with Elaborate Expression Annotations (2022.lrec-1)
Copied to clipboard
| Challenge: | SCH is the first substantial corpus to be annotated for elaborate expressions . a plurality of speakers are located in China, but many Hmong speakers left Laos as refugees . |
| Approach: | They describe the first publicly available corpus of Hmong, a minority language of China, Vietnam, Laos, Thailand, and various countries in Europe and the Americas. |
| Outcome: | The first publicly available corpus of Hmong is scraped from a long-running Usenet newsgroup . it is the first substantial corpus to be annotated for elaborate expressions . |
Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations (D18-1)
Copied to clipboard
| Challenge: | Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages . |
| Approach: | They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units. |
| Outcome: | The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training . |
Improved Neural Protoform Reconstruction via Reflex Prediction (2024.lrec-main)
Copied to clipboard
| Challenge: | comparative method allows linguists to infer protoforms from their reflexes based on sound change . authors argue that this approach ignores one of the most important aspects of the comparative approach . |
| Approach: | They propose a comparative method that allows linguists to infer protoforms from their reflexes . they propose to use a system where candidate protoform from a reconstruction model are reranked by a reflex prediction model. |
| Outcome: | The comparative method surpasses state-of-the-art methods on Chinese and Romance datasets. |
Cross-Cultural Similarity Features for Cross-Lingual Transfer Learning of Pragmatically Motivated Tasks (2021.eacl-main)
Copied to clipboard
| Challenge: | a large amount of work on cross-lingual transfer learning focused on typological and genealogical similarities between languages. |
| Approach: | They propose three features that capture cross-cultural similarities that manifest in linguistic patterns and quantify distinct aspects of language pragmatics. |
| Outcome: | The proposed features capture cross-cultural similarities manifest in linguistic patterns and quantify aspects of language pragmatics. |
ZIPA: A family of efficient models for multilingual phone recognition (2025.acl-long)
Copied to clipboard
| Challenge: | IPA transcriptions capture major articulatory contrasts in speech sounds, including the voicing status, place of articulation, manner of voicing, and tongue positions. |
| Approach: | They present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. |
| Outcome: | The proposed model outperforms existing phone recognition systems on 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. |
WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)
Copied to clipboard
| Challenge: | Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count. |
| Approach: | They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP. |
| Outcome: | The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task . |
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model (2026.acl-long)
Copied to clipboard
Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David R. Mortensen, Shinji Watanabe
| Challenge: | Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets. |
| Approach: | They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together . |
| Outcome: | The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR. |
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration (2026.findings-acl)
Copied to clipboard
Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank
| Challenge: | We show that script information is linearly encoded in the activation space of multilingual speech models . modifying activations at inference time induces script change even in unconventional pairings . |
| Approach: | They propose to add script vectors to activations at test time to induce script change . they also show that script information is linearly encoded in the activation space of multilingual speech models . |
| Outcome: | The proposed approach can induce script change even in unconventional language-script pairings. |
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs (2024.lrec-main)
Copied to clipboard
| Challenge: | Lexical-syntactic flexibility is a hallmark of English morphology . conversion involves placing a word with one part of speech in a non-prototypical context . |
| Approach: | They propose to test lexical-syntactic flexibility in the form of conversion . conversion is a process where a word with one part of speech is placed in a non-prototypical context . |
| Outcome: | The proposed task tests the ability of five language models to generalize over words with a non-prototypical part of speech. |
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods (2026.eacl-long)
Copied to clipboard
| Challenge: | a key problem in multilingual NLP is script barrier, which makes it difficult to share knowledge between languages . a new study shows that transliteration can be useful for languages using non-Latin scripts . |
| Approach: | They propose to use romanization, phonemic transcription, and substitution ciphers to evaluate models . romanization outperforms other input types in 7 out of 8 evaluation settings . |
| Outcome: | The proposed approach outperforms other input types on three tasks and is the most effective . romanization outperformed other input type in 7 out of 8 evaluation settings . |
Programming by Example meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction (2025.acl-long)
Copied to clipboard
Atharva Naik, Darsh Agrawal, Hong Sng, Clayton Marr, Kexun Zhang, Nathaniel Romney Robinson, Kalvin Chang, Rebecca Byrnes, Aravind Mysore, Carolyn Rose, David R. Mortensen
| Challenge: | Historical linguists have written programs that convert reconstructed words into their attested descendants via ordered string rewrite functions. |
| Approach: | They propose to use a model to generate a "similar distribution" for sound law induction . they propose four kinds of methods with varying amounts of inductive bias to investigate best performance . |
| Outcome: | The proposed model shows that it can be fine tuned with training data and evaluation data. |
Automatic Extraction of Rules Governing Morphological Agreement (2020.emnlp-main)
Copied to clipboard
Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig
| Challenge: | Creating a descriptive grammar is an indispensable step for language documentation but it is tedious and time-consuming. |
| Approach: | They propose a framework for extracting a first-pass grammatical specification from raw text in a concise, human- and machine-readable format. |
| Outcome: | The proposed framework extracts a grammatical specification that is nearly equivalent to those created with large amounts of gold-standard annotated data. |
Phonotactic Complexity across Dialects (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent studies show a moderate negative correlation between phonotactic complexity and word length in 106 languages. |
| Approach: | They propose to use a phone-level language model to measure phonotactic complexity . they find a tradeoff between word length and phonomactic complex . |
| Outcome: | The proposed model shows that low phonotactic complexity dialects concentrate around capital regions. |
Epitran: Precision G2P for Many Languages (L18-1)
Copied to clipboard
| Challenge: | Epitran is a multilingual, multi-back-end system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under an MIT license . |
| Approach: | Epitran is a multilingual back-end system for grapheme-to-phoneme transduction . it takes word tokens in the orthography of a language and outputs a phonemic representation . Epitran's efficacy has been demonstrated in multiple research projects . |
| Outcome: | Epitran is a multilingual, multi-backend system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under MIT license . |
Searching for the Most Human-like Emergent Language (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on emergent communication systems to generate languages with high statistical similarity to human languages has not been done. |
| Approach: | They propose to optimize a signalling game-based emergent communication environment to generate state-of-the-art emergentic languages with a high degree of similarity to human language. |
| Outcome: | The proposed language generates state-of-the-art on XferBench benchmark, demonstrating its similarity to human language and entropy-minimization properties. |
[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on how self-supervised speech models encode rich phonetic information have not explored how they are structured. |
| Approach: | They conduct a comprehensive analysis of the underlying structure of S3M representations with particular attention to phonological vectors. |
| Outcome: | The proposed model encodes phonologically interpretable and compositional vectors, demonstrating phonology vector arithmetic. |
PRiSM: Benchmarking Phone Realization in Speech Models (2026.acl-long)
Copied to clipboard
Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi, Eunjung Yeo, Ryan Soh-Eun Shim, Hanyu Zhou, Brendon Boldt, Karen Rosero, Kalvin Chang, Darsh Agrawal, Keer Xu, Chao-Han Huck Yang, Jian Zhu, Shinji Watanabe, David R. Mortensen
| Challenge: | Existing evaluations of phone recognition systems only measure surface-level transcription accuracy. |
| Approach: | They propose to standardize transcription-based evaluation and assess downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. |
| Outcome: | The proposed system outperforms LALMs in clinical, educational, and multilingual settings. |
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate (2024.lrec-main)
Copied to clipboard
Vilém Zouhar, Kalvin Chang, Chenxuan Cui, Nate B. Carlson, Nathaniel Romney Robinson, Mrinmaya Sachan, David R. Mortensen
| Challenge: | Existing word embedding methods overlook phonetic information that is crucial for many tasks. |
| Approach: | They propose three methods that use articulatory features to build phonetically informed word embeddings. |
| Outcome: | The proposed methods improve word retrieval and correlation with sound similarity and on rhyme and cognate detection tasks. |
Evaluating the Morphosyntactic Well-formedness of Generated Texts (2021.emnlp-main)
Copied to clipboard
Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, Yulia Tsvetkov
| Challenge: | Text generation systems are ubiquitous in natural language processing applications, but evaluation of these systems remains a challenge, especially in multilingual settings. |
| Approach: | They propose a metric to evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language. |
| Outcome: | The proposed metric can evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language. |
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics (2026.findings-acl)
Copied to clipboard
Atharva Naik, null Prakam, Yash Mathur, Darsh Agrawal, Manav Nitin Kapadnis, Yuwei An, Clayton Marr, Carolyn Rose, David R. Mortensen
| Challenge: | a benchmark for inductive reasoning is based on sound law induction in historical linguistics . solve rates are below 5% on hard PBEBench instances with long program cascades despite expensive scaling strategies . |
| Approach: | They propose a benchmark for inductive reasoning inspired by sound law induction in historical linguistics. |
| Outcome: | The proposed approach generates problems with controllable difficulty and ordering constraints . solve rates remain below 5% on hard PBEBench instances with long program cascades . |
Morpheme Induction for Emergent Language (2025.emnlp-main)
Copied to clipboard
| Challenge: | CSAR is a greedy algorithm that weights morphemes based on mutual information between forms and meanings, then removes it from the corpus and repeats the process to induce more morphs. |
| Approach: | They propose an algorithm that weights morphemes based on mutual information between forms and meanings, selects highest-weighted pair, removes it from corpus, and repeats process to induce further morphs. |
| Outcome: | The proposed algorithm makes reasonable predictions in adjacent domains. |
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models (2025.acl-long)
Copied to clipboard
Niyati Bafna, Emily Chang, Nathaniel Romney Robinson, David R. Mortensen, Kenton Murray, David Yarowsky, Hale Sirin
| Challenge: | Recent advances in MT quality and language coverage have shown that language varieties with low baseline performance are more likely to benefit from these approaches. |
| Approach: | They propose a training-time technique for adapting a pretrained model to dialectal data and an inference-time intervention adapting dialectal datasets to the model expertise. |
| Outcome: | The proposed model shows significant performance gains for several dialects from four language families, and modest gains for two other language families. |
Phone Inventories and Recognition for Every Language (2022.lrec-1)
Copied to clipboard
| Challenge: | Identifying phone inventories is crucial component in language documentation and preservation of endangered languages. |
| Approach: | They propose a probabilistic and non-probabilistic phone inventory model that estimates the phone inventory for any language listed in Glottolog. |
| Outcome: | The proposed model outperforms baseline models by 6.5 F1 and improves the PER (phone error rate) in phone recognition by 25%. |
Transformed Protoform Reconstruction (2023.acl-short)
Copied to clipboard
| Challenge: | Historical linguists reconstruct proto-languages by identifying systematic sound changes that can be inferred from correspondences between attested daughter languages. |
| Approach: | They propose to update their Latin protoform reconstruction model with the Transformer . romance data of 8,000 cognates spanning 5 languages and Chinese dataset are outperformed . |
| Outcome: | The proposed model outperforms previous models on Romance and Chinese datasets. |
Communicating in Emergent Language with an Induced Morphological Phrasebook (2026.acl-long)
Copied to clipboard
| Challenge: | a major challenge in studying emergent languages is interpreting how they convey meaning-neural networks may invent communication systems lacking features of human language. |
| Approach: | They build rule-based emergent language agents using form-meaning mappings induced from ELs and test their communicative performance in the EL environment. |
| Outcome: | The proposed model shows that EL agents rely on repetition and morpheme ordering to convey meaning. |
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons (2024.lrec-main)
Copied to clipboard
| Challenge: | In this paper, we examine the ability of large language models (LLMs) to identify different meanings in sentences that are superficially similar. |
| Approach: | They propose a challenge dataset for NLP with large lexical overlap which minimises the possibility of models discerning entailment solely based on token distinctions. |
| Outcome: | The proposed model fails to distinguish between constructions with three classes of adjectives which cannot be distinguished by surface features. |
AlloVera: A Multilingual Allophone Database (2020.lrec-1)
Copied to clipboard
David R. Mortensen, Xinjian Li, Patrick Littell, Alexis Michaud, Shruti Rijhwani, Antonios Anastasopoulos, Alan W Black, Florian Metze, Graham Neubig
| Challenge: | Phonemes are contrastive phonological units, and allophones are their various concrete realizations. |
| Approach: | They propose a resource that maps allophones to phonemes for 14 languages . they propose phonological representations that are much closer to a universal transcription . |
| Outcome: | The proposed resource maps from 218 allophones to phonemes for 14 languages. |